Papers with sequence model
A Multiscale Visualization of Attention in the Transformer Model (P19-3)
Copied to clipboard
| Challenge: | Various tools have been developed to visualize attention in NLP models, ranging from attention-matrix heatmaps to bipartite graph representations. |
| Approach: | They propose an open-source tool that visualizes attention at multiple scales and provides a unique perspective on the attention mechanism. |
| Outcome: | The proposed model outperforms OpenAI GPT-2 and BERT on several language modeling benchmarks. |
Self-Attentive Residual Decoder for Neural Machine Translation (N18-1)
Copied to clipboard
| Challenge: | Neural sequence-to-sequence networks with attention have been used for machine translation . however, the target-side context is limited and the model lacks the ability to capture non-syntactic dependencies among words. |
| Approach: | They propose a sequence-to-sequence network with attention that captures contextual information at each time-step prediction through an attention mechanism. |
| Outcome: | The proposed model outperforms a neural MT baseline and memory and self-attention network on three language pairs. |
Searchable Hidden Intermediates for End-to-End Models of Decomposable Sequence Tasks (2021.naacl-main)
Copied to clipboard
| Challenge: | ESPnet framework exploits compositionality to learn searchable hidden representations at intermediate stages of a sequence model using decomposed sub-tasks. |
| Approach: | They propose a framework that exploits compositionality to learn searchable hidden representations at intermediate stages of a sequence model using decomposed sub-tasks. |
| Outcome: | The proposed framework outperforms the state-of-the-art on speech translation tasks by +6 and +3 BLEU on the two test sets of Fisher-CallHome and +4 BLUE on the English-German and English-French test sets. |
Determinantal Beam Search (2021.acl-long)
Copied to clipboard
| Challenge: | a new beam search approach allows for a diverse subset selection process . standard beam search does not encode interactions between candidates . |
| Approach: | They propose a beam search reformulation that casts subset selection as the subdeterminant optimization problem. |
| Outcome: | The proposed method offers competitive performance against other diverse set generation strategies while providing a more general approach to optimizing for diversity. |
The Unstoppable Rise of Computational Linguistics in Deep Learning (2020.acl-main)
Copied to clipboard
| Challenge: | a quarter century ago, linguists assumed that language knowledge needed to be innate . but vector-space representations and machine learning algorithms are much more powerful than was thought . |
| Approach: | They trace the history of neural networks applied to natural language understanding tasks . they argue that Transformer is not a sequence model but an induced-structure model . |
| Outcome: | The proposed model is not a sequence model but an induced-structure model, the authors argue . they argue that the nature of language has had a profound impact on progress in machine learning . |
Parsing as Tagging (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for dependency parsing treat parse as tagging, but they are not perfect. |
| Approach: | They propose a simple yet accurate method that treats parsing as tagging . they use a sequence model with a bidirectional LSTM over BERT embeddings . |
| Outcome: | The proposed method outperforms the state-of-the-art method on universal dependency (UD) by 1.76% unlabeled attachment score (UAS) for English, 1.98% UAS for French, and 1.16% UAS in German. |